Papers with multimodal translation systems
Understanding the Effect of Textual Adversaries in Multimodal Machine Translation (D19-64)
Copied to clipboard
| Challenge: | Existing studies show that multimodal machine translation systems are better than text-only systems at translating phrases that have a direct correspondence in the image. |
| Approach: | They conduct experiments with both visual and textual adversaries to understand the role of textual inputs in multimodal machine translation. |
| Outcome: | The proposed model can recover masked tokens in the source sentences . the proposed model is based on a model with a visual modality . |
Adversarial Evaluation of Multimodal Machine Translation (D18-1)
Copied to clipboard
| Challenge: | Existing evidence that visual context helps multimodal machine translation systems is unconvincing due to inconsistencies between text-similarity metrics and human judgements. |
| Approach: | They propose an adversarial evaluation method to examine the utility of image data in multimodal machine translation. |
| Outcome: | The proposed method shows that only one out of three publicly available systems is sensitive to this perturbation of the data. |
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation (2025.findings-emnlp)
Copied to clipboard
Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Vladimir Araujo, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofía Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio
| Challenge: | a human-curated benchmark of over 5,800 triples of images is used to evaluate multimodal translation systems. |
| Approach: | They introduce a human-curated benchmark of over 5,800 triples of images along with parallel captions in English and regional languages. |
| Outcome: | The results show that visual context improves translation quality in culturally-specific items . |